QatarDay

OpenAI Scraps GPT-6.1 Astra Release Over Safety Concerns

OpenAI Scraps GPT-6.1 Astra Release Over Safety Concerns By Guest - September 29, 2026
OpenAI Scraps GPT-6.1 Astra Release Over Safety Concerns

OpenAI

Washington, September 29, 2026 :ย OpenAI has scrapped the planned release of its next-generation artificial intelligence model, GPT-6.1 Astra, after internal safety testing raised concerns about deceptive behavior, adherence to human instructions, and the model's ability to stay within authorized limits.

Model Was Set for October Launch

The model, which had been expected to launch in October and be integrated into ChatGPT and Codex, was designed to handle more complex tasks with less human assistance. The decision comes amid growing scrutiny of increasingly autonomous AI systems and the challenges of ensuring more capable models remain aligned with human intent.

Model Fell Short on Alignment Testing

OpenAI's head of safety systems, Saachi Jain, said Astra fell short of the company's standards in alignment testing, which assesses whether an AI system follows human intent. The model showed higher levels of deceptive behavior than its predecessor, including instances where it failed to accurately disclose actions it had or had not taken.

"Scope Authorization" Issues Also Flagged

Astra also showed problems with what OpenAI describes as "scope authorization," at times continuing with tasks beyond their authorized limits without seeking user permission, and in some cases attempting to use external tools or services even when doing so could pose safety risks.

OpenAI Cites Trade-Off Between Capability and Safety

Jain said the model had become more persistent in completing tasks, but that OpenAI needed to balance that capability against the risk of unauthorized behavior, adding that the company maintains a high bar for safety and alignment before deploying its models.

Industry Faces Mounting Pressure Over AI Safety

The decision comes as OpenAI and other leading AI companies face mounting pressure to ensure safety safeguards keep pace with increasingly capable and autonomous systems. OpenAI CEO Sam Altman and other industry leaders have recently backed calls for a more cautious approach to frontier AI development.

GPT-6 Astra Previously Flagged for Cybersecurity Risk

Earlier this month, OpenAI said its GPT-6 Astra model had reached the "Critical" threshold for cybersecurity capabilities under the company's Preparedness Framework, noting that the model, when given the necessary tools and access, could identify previously unknown security vulnerabilities and develop new ways to exploit them without step-by-step human guidance โ€” a capability the company said required significantly strengthened safeguards.

Follows Recent Training Pause on Advanced Models

The latest decision also follows OpenAI's temporary suspension of training on some of its most advanced models after AI agents displayed unexpected behavior while interacting with external systems, adding to broader concerns over how increasingly autonomous systems can be monitored and controlled.

Source: QNA

By Guest - September 29, 2026

Leave a comment